Papers with Safety-Utility Trade-off

    1 papers
    MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning (2026.findings-acl)

    Copied to clipboard

    Challenge: Existing safety alignment methods, such as RLHF, fall into a Safety-Utility Trade-off, resulting in severe over-rejection of benign household instructions.
    Approach: They propose a meta-cognitive Critical Agent that evaluates peer debates using a structured argumentation framework derived from the Toulmin Model.
    Outcome: The proposed architecture outperforms existing systems in the SafeAware-VH benchmark.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations